NOTE

1.1 Elasticsearch translog

English translation of the original VNote ‘Elasticsearch translog’, preserving its structure and historical learning notes.

Elasticsearch / SearchCreated Updated 1 min readhistorical

This is a historical learning note and may contain outdated or incomplete understanding.

1. What Is translog?

When Elasticsearch performs a write operation (index or delete), it writes that operation to a log.

2. Why Is translog Needed?

  • Prevent data loss caused by a crash.
    • After Elasticsearch data is refreshed into the filesystem cache, it becomes searchable. If the system crashes at this point, the data may be lost, but performing a flush to write it to disk is too slow.
    • The translog is a sequentially written disk log, so it is fast.
    • Note that the translog also starts in memory. Only after fsync (not flush) writes it to disk can data loss be prevented.
  • Provide real-time CRUD.
    • When retrieving, updating, or deleting a document by ID, Elasticsearch first checks whether there are any recent changes in the translog before trying to retrieve the document from the relevant segment. This means the latest known version of the document can always be accessed in real time.

3. When translog fsync Happens

  • If index.translog.durability is request (the default), every write operation is fsynced to disk.
  • If index.translog.durability is async, fsync happens only when one of the following conditions is reached:
    • The translog size reaches index.translog.flush_threshold_size, default 512 MB.
    • The time reaches index.translog.sync_interval, default 5s.

4. References

Discussion

Sign in with GitHub to comment. Discussions are stored as GitHub Issues.View on GitHub